Back

Journal of the American Medical Informatics Association

Oxford University Press (OUP)

Preprints posted in the last 90 days, ranked by how well they match Journal of the American Medical Informatics Association's content profile, based on 71 papers previously published here. The average preprint has a 0.14% match score for this journal, so anything above that is already an above-average fit.

1
Predicting Deprescribing of High-Risk Medications Using Provider EHR Use and Patient Characteristics

Gu, B.; Jungo, K. T.; Lauffenburger, J.; Choudhry, N.; Isaac, T.; Zambrano, J.; Yang, J.

2026-07-29 health informatics 10.64898/2026.07.25.26358746 medRxiv
Top 0.1%
59.4%
Show abstract

Potentially inappropriate medications expose older adults to preventable harm, yet deprescribing remains difficult to implement consistently. Although electronic health record (EHR) interventions can reduce prescribing, health systems lack clear evidence about which routinely captured patient, primary care provider (PCP), and intervention-design factors predict medication discontinuation or dose tapering. Understanding these determinants is essential for targeting and scaling deprescribing support. In this study, we conducted the first machine-learning analysis of these trial data. We analyzed 2,979 adults aged 65 years or older and 158 structured EHR features spanning patient characteristics, PCP characteristics and EHR-use behaviors, and deprescribing-tool design. We compared eight models for predicting medication discontinuation or dose tapering and used SHAP to examine feature importance. TabPFN achieved the highest positive predictive value (67.71%), AUROC (74.30%), and AUPRC (60.95%), although overall predictability was moderate. Our findings show that even rich structured EHR data only moderately predict deprescribing, suggesting that important clinical determinants are not captured in routine fields. PCP EHR-use measures accounted for 19 of the 25 highest-ranked TabPFN features, although they also constituted most candidate predictors. The study provides the health system with an informative reference for predicting high-risk medication deprescribing. Future models should incorporate richer clinical context and undergo external validation before informing personalized deprescribing support.

2
The EHR Density Index: A new method to control for EHR data inconsistency across patients

Bhatia, A.; Lash, S.; McIntee, T.; Pfaff, E.

2026-08-06 health informatics 10.64898/2026.08.03.26359595 medRxiv
Top 0.1%
51.9%
Show abstract

Electronic health record (EHR) data vary substantially in documentation density across patients, independent of disease burden. Existing tools such as the Charlson Comorbidity Index (CCI) and Elixhauser Comorbidity Index measure disease burden but do not capture differences in data volume, leaving a common source of bias unaddressed in EHR-based analyses. To address this gap, we developed the EHR Density Index (EDI), which characterizes the quantity, depth, and breadth of EHR data per patient per year, normalized by utilization patterns, using records from 24,987 adult patients at UNC Health (2018 - 2024). The EDI combines a utilization cluster assigned via Gaussian Mixture Model with within-cluster residuals quantifying documentation volume across four clinical domains. Four interpretable clusters emerged; while CCI predicted cluster membership, its associations with within-cluster residuals were weak, confirming the EDI captures dimensions of the patient record distinct from disease burden. The EDI is intended as a covariate to address documentation density as a source of confounding in real-world data-driven research.

3
Sharing Aggregated Patient Counts in Place of Line-Level EHR Data: Analytic Fidelity and the Limits of Count Suppression for Privacy

Chen, Y.; McMurry, A.; Gottlieb, D.; Jones, J. R.; Strober, B. J.; Mandl, K. D.

2026-08-21 health informatics 10.64898/2026.08.18.26359984 medRxiv
Top 0.1%
51.2%
Show abstract

Objective. Privacy regulation constrains sharing line-level electronic health records (EHR) across institutions. One alternative is to aggregate counts into a cube, a table of counts for every combination of categorical variables, with cells below a threshold suppressed. This study asked whether common analyses on the cube reproduce conclusions from line-level data, and whether suppression prevents recovery of the small cells it is meant to hide. Materials and Methods. A Bayesian count-inference pipeline was built that reconstructs suppressed counts and doubles as a reconstruction attack. Applied to 285 pediatric kidney-transplant patients at Boston Children's Hospital, statistical fidelity (Jensen-Shannon divergence, Cramer's V, and R2) and analytical utility (marginal distributions, subgroup graft rejection odds ratios, and logistic-regression classification) were evaluated. Conditional Tabular GAN (CTGAN) synthetic data served as a comparator. Results. Statistical analyses on the cube recapitulated results from line-level data. Across 106 demographic-by-medication subgroups, a bootstrap mean of 3.5 subgroups showed a significant graft-rejection association. The cube's odds-ratio sign changes reversed no significant associations, versus 2.3 for CTGAN. The same reconstruction also defeated suppression: in a 10-variable cube, 76.6% of suppressed cube cells were recovered exactly (14,554 of 18,994), including 85.5% of single-patient cells. Discussion. The cube reproduced common kidney-transplant analyses, but the same reconstruction also recovered suppressed cells; fidelity and privacy risk are thus two faces of one reconstruction rather than independent properties. Conclusions. The cube is a useful surrogate for these kidney-transplant analyses only when paired with a stronger privacy mechanism. This study demonstrated reconstructability of suppressed counts, not re-identification.

4
REFINE: Closing the Loop Between Large Language Models and Symbolic Rules in Clinical NLP

Wang, N.; Kakadiaris, A.; Li, C.; Wang, R.; Ahn, J.; Wang, Y.; Fu, S.

2026-08-17 health informatics 10.64898/2026.08.11.26360118 medRxiv
Top 0.1%
45.4%
Show abstract

Symbolic clinical natural language processing (NLP) systems remain widely used for extracting clinical concepts from electronic health record (EHR) narratives, but maintaining rule resources requires extensive manual error analysis and rule refinement. This study investigates whether large language models (LLMs) can assist in identifying extraction errors and generating candidate rules to improve symbolic clinical NLP systems. Using error reports derived from a multi-site evaluation of a previously validated symbolic model for cognitive and neuropsychiatric-related clinical concepts, we developed a human-in-the-loop framework, REFINE. The framework first uses LLMs to classify extraction errors and generate explanatory reasoning, which can then be incorporated into prompts for rule generation. Three LLMs (GPT-5.2, GPT-4o, GPT-4o-mini) were evaluated under four prompting conditions. LLM-generated rule sets improved performance compared with the baseline NLP-CAM system, increasing F1-score from 0.37 to 0.58. These findings suggest that LLMs can support scalable rule refinement for symbolic clinical NLP systems.

5
A PRISMA-Aligned Agentic Framework for Medical Systematic Reviews and Evidence Synthesis

Huang, H.; Zheng, Q.; Qiu, P.; Zhao, W.; Zhang, Y.; Xie, W.; Wang, Y.; Zhang, X.; Wu, C.

2026-08-02 health informatics 10.64898/2026.07.30.26359375 medRxiv
Top 0.1%
45.2%
Show abstract

Medical systematic reviews are central to evidence-based medicine, but they remain slow, labor-intensive, and difficult to maintain under the full Preferred Reporting Items for Systematic Reviews and Meta-Analyses (PRISMA) workflow. Recent LLM-based deep research agents offer a promising route to addressing this challenge, yet reliable deployment in medical systematic reviews remains limited by insufficient clinical domain knowledge and inconsistent adherence to evidence-based methodological standards across the full workflow. We address these gaps with MedSR-Copilot, a PRISMA-aligned multi-agent copilot that decomposes review automation into literature retrieval, coarse-to-fine screening, data extraction, Risk-of-Bias assessment, and evidence synthesis, while preserving structured intermediate artifacts throughout the workflow. We further introduce MedSR-Bench, an end-to-end benchmark for evaluating systems beyond isolated subtasks, from review input to final evidence-synthesis conclusions. MedSR-Copilot completes medical systematic reviews end-to-end under the full PRISMA workflow, achieving 63.6% human-aligned conclusions, 18.3 percentage points above the best baseline among strong general-purpose LLMs and prior automated review systems. In a human-AI collaboration study involving 23 analysis groups across four systematic review topics, MedSR-Copilot, used as a copilot, reduces end-to-end review time by 64.9% and improves final conclusion accuracy by 27.4 percentage points compared with routine-practice workflows. Together, these results demonstrate the reliability and efficiency of MedSR-Copilot as a medical research copilot and suggest a practical path toward trustworthy review automation.

6
A language model framework for sequence modeling of EHR audit logs to characterize clinician-EHR interactions

Kim, S.; Lou, S. S.; Cobb, A.; Jha, S.; Kannampallil, T. G.

2026-06-26 health informatics 10.64898/2026.06.24.26356449 medRxiv
Top 0.1%
39.0%
Show abstract

Objective: Electronic health record (EHR) audit logs capture clinician-EHR interaction patterns, but most audit log research relies on aggregated measures (e.g., total time). We investigated how audit logs could be modeled using large language model (LLM) architectures to learn underlying workflow sequences during clinician-EHR interactions. Materials and Methods: Using >295 million EHR-based audit log actions from inpatient settings spanning 2019 to 2024, we fine-tuned Llama-3-8B under two encodings: (1) symbolic field-based tokens and (2) semantic natural language audit log action descriptions. A first-order Markov model, which used only the immediate prior action, served as baseline minimal context comparator. Model representation was assessed using next-action prediction accuracy in an early-period test set and two temporally distinct out-of-sample (OOS) periods. Results: In the early period test set, the semantic LLM achieved the highest accuracy (0.7418, 95%CI [0.7415 - 0.7420]) compared to the symbolic LLM (0.3838, 95% CI: 0.3836 - 0.3840) and Markov baseline (0.4553, 95%CI [0.4551 - 0.4555]). The semantic approach was also robust to temporal drift in EHR interaction patterns seen in the two OOS periods (semantic LLM vs Markov accuracy, OOS-1: 0.6509 vs 0.3169; OOS-2: 0.6232 vs 0.2648). Discussion and Conclusions: Semantic LLM relying on audit log action descriptions yielded the highest next-action prediction accuracy, and demonstrated robustness to temporal drift, suggesting that longer sequential context and semantic action descriptions may improve audit log-based sequence modeling. These findings support further development of semantic sequence models for audit log research, including task identification, automated workflow characterization, and safety-focused analyses of clinician-EHR interaction patterns.

7
SPIRIT-CONSORT-ELM: Element-Level Assessment of Randomized Controlled Trial Reporting Using Large Language Models

Jiang, L.; Ying, X.; Brown, A. W.; Lan, M.; Song, W.; Menke, J.; Vorland, C.; Mayo-Wilson, E.; Kilicoglu, H.

2026-06-15 health informatics 10.64898/2026.06.06.26354746 medRxiv
Top 0.1%
38.8%
Show abstract

Randomized controlled trials (RCTs) play a central role in assessing the benefits and harms of interventions. Incomplete reporting in RCT publications can compromise the verifiability and usefulness of RCTs. SPIRIT and CONSORT reporting guidelines aim to improve the completeness of RCT protocols and results publications, respectively. However, many RCTs are not reported completely. Checking manuscripts automatically could help authors improve the completeness of reports prior to publication. We previously annotated SPIRIT-CONSORT-TM, a corpus of 200 articles (comprising 100 protocol-results publication pairs) using 83 checklist items drawn from SPIRIT 2013 and CONSORT 2010. We also trained machine learning models to automatically assess reporting at the item level. Each checklist item can include multiple constituent elements (i.e., specific details required for that item), and an item might be considered fully reported when all of its elements are present. However, prior work does not explicitly capture or evaluate reporting at the element level. To address this gap, we extended SPIRIT-CONSORT-TM by incorporating element-level annotations and using them to assess reporting completeness (SPIRIT-CONSORT-ELM). We formulated element-level assessment as a machine reading comprehension task, operationalized through 119 questions, where each question targets a specific reporting element within a checklist item. Using the 200 articles included in SPIRIT-CONSORT-TM, two annotators independently answered 119 questions for 50 articles (25 protocol-results pairs) and resolved any discrepancies through discussion; the remaining 150 articles (75 protocol-results pairs) were assessed by a single annotator. We then developed an automated pipeline for element-level assessment using SPIRIT-CONSORT-ELM. The pipeline first applies a PubMedBERT-based model to identify sentences containing item-level reporting information, then it uses a generative large language model (LLM; GPT-5) with chain-of-thought reasoning to answer element-level questions based on the retrieved evidence. Agreement between the two annotators was high (Gwet's AC1: 0.782) and our pipeline achieved high accuracy in identifying element-level reporting evidence (F1: 0.822, Gwet's AC1: 0.796). Ablation studies indicate that chain-of-thought reasoning and the inclusion of illustrative in-context examples modestly improve LLM performance on the machine reading comprehension task. SPIRIT-CONSORT-ELM provides a benchmark for evaluating reporting guideline completeness at the element level, enabling assessment of RCT transparency beyond the simple presence or absence of checklist items and is publicly available at https://osf.io/kznx4/. The automated pipeline establishes a robust baseline for assessing RCT reporting and demonstrates potential as a practical aid for authors, reviewers, and editors to identify and address gaps in completeness and transparency of RCT reports.

8
Evaluating Clinical Concept Extraction and Evidence-Bounded Terminology Linking: Multisite Model Comparison and Pilot Ablation Study

Chen, Y.; Popescu, M.

2026-08-24 health informatics 10.64898/2026.08.20.26360740 medRxiv
Top 0.1%
38.8%
Show abstract

Background: Clinical terminology pipelines must first extract candidate spans from narrative notes and then determine whether those spans map to existing concepts or warrant further review. Evaluation is difficult because span boundaries vary between annotators and because downstream decisions depend on the terminology evidence retrieved for each span. Objective: We evaluated clinical concept extraction, terminology linking across controlled evidence conditions, and ontology-extension triage for terms that remained unmatched after initial terminology screening. Methods: We conducted 3 complementary pilot evaluations that used distinct units of analysis and were analyzed separately. Study 1 compared 5 automated extraction pipelines and a union-merge analysis with 2 human annotation sets in 66 deidentified clinical notes from 3 health systems. Agreement was evaluated by exact string matching and BGE-large-en-v1.5 embedding matching. Study 2 evaluated 56 clinical spans, including 28 with reference Unified Medical Language System concepts and 28 adjudicated as unsuitable for ontology extension, under complete retrieval, matched-concept masking, and large language model-only inference, yielding 168 span-condition outputs. The graph retrieval pipeline used BGE-large-en-v1.5 embeddings, and the decision model was Gemma 3 27B. Study 3 applied full vector retrieval to 84 terms previously not matched in either UMLS or BioPortal. Results: In Study 1, interannotator exact-match F1 was 0.29 and embedding-match F1 was 0.75. Automated exact-match F1 scores ranged from 0.07 to 0.17; embedding-match F1 was highest for MedGemma (0.55), followed by Gemma (0.53), sci_md and SciBERT (each 0.43), and Llama 3.3 (0.32). In Study 2, complete retrieval returned a reference-matched link for 28/28 known-concept spans (100%; 95% CI, 87.9%-100%). Masking assigned POSSIBLE_CANDIDATES to all 28; large language model-only inference assigned POSSIBLE_CANDIDATES to 25/28 (89.3%) and LINKED to 3/28 (10.7%). Across the 3 evidence conditions, the same 12/28 unsuitable-extension spans were classified as NOT_MEANINGFUL (42.9%) and the same 16/28 as POSSIBLE_CANDIDATES (57.1%). In Study 3, the pipeline assigned PLAUSIBLE_EXISTING_CONCEPT to all 84 terms, none was flagged for extension, and top-candidate similarity averaged 0.914 (SD 0.027); extension status was not independently adjudicated. Conclusions: Measured extraction performance varied substantially by matching definition, whereas exact-link decisions varied with the availability of matched terminology evidence. In the follow-up sample, initial nonmatching did not establish ontology novelty: after semantic retrieval, the pipeline classified all 84 terms as plausible existing concepts and proposed none for extension. These findings support separate evaluation of extraction, retrieval, evidence-grounded linking, and extension candidacy.

9
Same Inputs, Different EDSS: Measuring Specification Drift in Clinical Scoring Pipelines

Hwang, S.; Mowery, D. L.; Thomas, S.; Williams, H.; Bar-Or, A.; Sharma, V.; Buijs, F.; Perrone, C.

2026-07-07 health informatics 10.64898/2026.06.25.26356350 medRxiv
Top 0.1%
38.7%
Show abstract

Clinical informatics pipelines increasingly compute validated clinical endpoints from upstream NLP outputs. Even when the endpoint is defined by an established rubric, translating that rubric across representations - natural language instructions, program logic, and reference implementations - can introduce specification drift, where ostensibly equivalent calculators yield meaningfully different scores. We study this phenomenon for the Expanded Disability Status Scale (EDSS), a standard measure of disability in multiple sclerosis. Holding constant a shared set of functional system (FS) subscores extracted by a large language model (LLM), we compare EDSS values computed across three representations of the same scoring rubric: prompt-executed natural language, LLM-generated code, and a canonical reference implementation. We characterize disagreement structure, distributional shifts, and clinically salient boundary flips, and we propose an audit workflow that treats endpoint computation as a first-class verification target in clinical NLP systems.

10
Geometry-Aware Reproducibility of Imputation Protocol (GRIP): diagnosing two failure modes of single imputation for heavy-tailed, collinear variables in biomedical data

Park, C. S.-Y.

2026-07-14 health informatics 10.64898/2026.07.14.26358017 medRxiv
Top 0.1%
38.5%
Show abstract

Background: Imputation accuracy is typically summarized by a single figure computed from one set of random deletions. For the heavy-tailed, collinear variables common in biomedical data, this study asks whether that figure can be trusted, and introduces a diagnostic protocol that determines when it cannot. Methods: This paper introduces GRIP (Geometry-aware Reproducibility of Imputation Protocol), a three-step diagnostic that profiles each variable's geometry and collinearity, stress-tests reproducibility under MCAR, MAR, and MNAR missingness with fixed seeds, and classifies the failure mode. GRIP was demonstrated on 1,885 United States Centers for Medicare & Medicaid Services (CMS) home-health agencies (15 numeric variables), comparing missForest, votingForest, mean, and median imputation. A supplementary k-nearest-neighbour (k-NN) grid (20 bivariate lognormal parameter combinations; n = 500) verified the SILENT failure boundary across imputer types. The simulation component followed the ADEMP framework. Results: Geometry profiling prospectively flagged two variables combining extreme right-skew (skewness 38.6, 43.3) with near-collinearity (|r| = 0.96). Under MCAR and MAR, missForest normalized root mean squared error (NRMSE) ranged from below 1 to 152 across otherwise identical replications (SD 14-30); the difficulty ordering of the two variables reversed in 69% of replications. Under MNAR self-masking, instability vanished (SD = 0), yet only 14% of true extreme magnitudes were recovered. Both failure modes arise from one mechanism: instability requires an extreme value to be absent while its collinear partner remains observed. A k-NN proxy grid confirmed SILENT failure in 11 of 15 high-skew parameter combinations under MNAR, regardless of correlation level. Conclusions: For heavy-tailed, collinear variables, one imputation-accuracy number can mislead in two opposite ways. GRIP detects both before an imputer is committed and is provided as reusable, open-source R code.

11
Agentic Artificial Intelligence for Hospital Readmission Review: A Single-Center Blinded Evaluation and Exploratory Qualitative Analysis

Gensheimer, M. F.; Adhikari, R.; Parmer-Chow, C.; Liu, N.; Ma, S.; Shieh, L.

2026-06-22 health systems and quality improvement 10.64898/2026.06.17.26355917 medRxiv
Top 0.1%
38.5%
Show abstract

Background: Manual review of 30-day hospital readmissions can identify actionable quality and safety problems, but it is labor-intensive. We developed and evaluated an agentic AI workflow for evidence-grounded readmission review. Materials and methods: We studied adult patients with unplanned 30-day readmission after discharge from a medicine hospitalist service at a single academic health system. An AI agent using a large language model queried a database containing notes, encounters, procedures, laboratory results, and other clinical data, and completed the same structured readmission-review rubric used by physicians. In the primary comparative evaluation, 20 randomly selected readmissions from 2025 were each reviewed by two physicians and the AI system. Blinded physician evaluators rated review quality. After rubric refinement, the AI workflow was applied to 100 recent readmissions in an exploratory expanded-cohort analysis of recurring improvement opportunities. Results: In the primary comparative evaluation, the AI classified 9/20 readmissions (45%) as preventable, compared with 19/40 physician reviews (47.5%). Blinded overall quality ratings were similar for AI and physician reviews (4.35 vs. 4.20 on a 1-5 scale; mean difference 0.15, 95% CI -0.20 to 0.48; p=0.49), as were factuality/support and usefulness/actionability ratings. No AI hallucinations were identified during factuality review. Agreement on preventability and primary readmission category was low for both AI-human and human-human comparisons. The AI system cost $0.23 per chart; physician reviewers took a median of 15 minutes, corresponding to an estimated $42.43 per chart. In the exploratory expanded-cohort analysis, AI-assisted review identified recurring vulnerabilities in post-discharge follow-up plans, incomplete inpatient workups, medication-safety transitions, and indwelling-device transitions. Conclusions: Agentic AI produced readmission reviews with similar blinded quality ratings to physician reviews in this small single-center primary comparative evaluation and supported identification of recurring quality-improvement themes in the exploratory expanded-cohort analysis. Preventability judgments remained variable among both AI and physicians, underscoring the need for human oversight and prospective evaluation before operational use.

12
Large-Scale Psychiatric Concept Extraction from Electronic Health Records: A Comparative Study of Encoder-Based Language Models

Xue, X.; Frydman-Gani, C.; Arias, A.; Perez Vallejo, M.; Londono Martinez, J. D.; Valencia-Echeverry, J.; Castano, M.; Freimer, N. B.; Lopez-Jaramillo, C.; Olde Loohuis, L. M.

2026-08-23 health informatics 10.64898/2026.08.20.26360921 medRxiv
Top 0.1%
37.9%
Show abstract

Background: Free-text notes in electronic health records (EHRs) contain fine-grained psychiatric information that is essential for psychiatric research and clinical care, and often absent or under-recorded in structured codes alone. Clinical natural language processing (cNLP) can support extraction of this information from EHR notes, yet Spanish-language cNLP remains under-developed. Moreover, broad evaluations comparing multiple encoder-based language models across extensive, fine-grained psychiatric concept sets remain scarce, and it remains unclear how these models compare with traditional NLP (tNLP) systems and much larger generative large language models (LLMs). In addition, cross-site performance of fine-tuned models is rarely tested, and limited annotated training data remains a major challenge, especially for rare symptoms. Objectives: We aimed to advance scalable, global psychiatric cNLP by fine-tuning multiple encoder-based models with differing architectures and pre-training strategies for detecting fine-grained psychiatric concepts in Spanish EHRs. We further evaluated the impact of augmenting the fine-tuning data with precision-weighted weak labels for less-frequent concepts, and compared the performance of the encoder-based models to that of tNLP and a fine-tuned generative LLM trained on the same data. Finally, we evaluated model cross-site generalizability on an external EHR dataset. Methods: Three encoder-based models (BETO, XLM-RoBERTa-large, and bsc-bio-ehr-es) were fine-tuned on 1,642 clinician-annotated EHR documents from Colombia to detect 110 psychiatric concepts in Spanish text. To address the limited annotated examples available for less-frequent concepts, 12,000 additional documents were weakly-labeled for less-frequent concepts using tNLP, and incorporated into the fine-tuning data with labels weighted by pattern precision. Models were compared with tNLP and a generative LLM, and evaluated on an external EHR dataset from another psychiatric hospital in Colombia. Results: Encoder model performance varied substantially, with macro-F1 ranging from 0.64 to 0.81. BETO achieved the highest macro-F1 (0.81; median F1=0.88 [IQR=0.77-0.96]). Adding precision-weighted weak labels for less-frequent concepts improved BETO's overall macro-F1 to 0.83 and increased mean F1 for the 55 augmented concepts from 0.82 to 0.86. Under matched fine-tuning conditions, fine-tuned BETO and the tNLP method were equivalent in F1, whereas the LLM significantly outperformed BETO in F1. After weak-label augmentation, BETO significantly outperformed tNLP in F1 (PFDR<.001) and narrowed the performance gap with the LLM, although equivalence was not established. Lastly, fine-tuned BETO maintained reasonably strong performance on data from an external hospital not used for model fine-tuning (out-of-domain macro-F1=0.78). Conclusions: General-purpose pre-trained encoders had strong performance for psychiatric concept extraction from Spanish EHRs. Weak-label augmentation improved BETO's performance and strengthened results relative to a tNLP baseline, while reducing, but not eliminating, the performance gap with a much larger fine-tuned generative LLM. These findings highlight the utility of these relatively lightweight models for scalable, accurate and reproducible detection of psychiatric concepts in Spanish-language EHRs.

13
A Comparison of Manual and Automated Approaches to Developing Computable Algorithms for Identifying Acute Pancreatitis

Bann, M. A.; Carrell, D. S.; Gruber, S.; Heagerty, P. J.; Williamson, B. D.; Nelson, J. C.; Hazlehurst, B.; Felcher, A.; Nyongesa, D. B.; Slaughter, M. T.; Sapp, D. S.; Cronkite, D. J.; Ball, R.; Floyd, J. S.

2026-06-08 health informatics 10.64898/2026.06.05.26354934 medRxiv
Top 0.1%
36.2%
Show abstract

Objective: Clinical phenotyping methods that rely on clinical and informatics expertise can be time-intensive and costly. We tested both manual and highly automated approaches using electronic health record (EHR) data to identify an FDA Sentinel Initiative health outcome of interest, acute pancreatitis. Materials and Methods: We trained and evaluated machine learning algorithms using EHR data with two approaches: a custom approach that included manually curated features and trained on outcomes data validated with medical record review, and a highly automated approach that greatly simplifies and automates feature engineering and relies on low-cost silver-standard outcomes for model training. Results: Custom algorithms using manually curated structured claims data discriminated cases from non-cases with a high degree of accuracy (cv-AUC 0.89 [95%CI 0.84-0.94]); the inclusion of natural language processing (NLP)-derived covariates from clinical notes increased performance slightly (cv-AUC 0.91[95%CI 0.86-0.97]). The automated algorithm trained on the outcome count of diagnosis codes performed less well (AUC 0.80 [95% CI 0.75-0.85]) but improved using maximum lipase value as an outcome (AUC 0.88 [95% CI 0.84-0.92]). At a positive predictive value of 90%, the custom algorithm had a sensitivity of 92%, the automated algorithm trained on diagnosis code count had a sensitivity of 45%, and the automated algorithm trained on maximum lipase value had a sensitivity of 84%. However, a prediction rule derived by clinicians during chart review was nearly as accurate (maximum lipase value [&ge;] 3 times upper limit of normal; AUC 0.86, PPV 85%, sensitivity 92%). Discussion: Machine learning algorithms with manually curated structured data and NLP features trained on validated outcomes data successfully identified validated events. Use of an outcome in the automated model based on specific phenotype knowledge (maximum lipase value) allowed for performance similar to the custom model and with considerably less resources.

14
TrialCode Agent: LLM-Assisted Clinical Code-Set Construction for Trial Emulation

Habibdoust, A.; Sajjad, A.; Hernandez, D.; Patel, K.; Song, X.

2026-08-23 health informatics 10.64898/2026.08.20.26360962 medRxiv
Top 0.1%
35.9%
Show abstract

Objective Translating free-text clinical trial criteria into computable code sets is a valuable standardization practice that is necessary for producing reproducible real-world evidence studies but requires standardized interpretation across multiple clinical vocabularies. Methods We developed TrialCode Agent, a hybrid-large language model (LLM)-terminology verification agent that generates, formats, verifies, and expands candidate codes from free-text clinical criteria. The system supports ICD-9-CM diagnoses and procedures, ICD-10-CM, ICD-10-PCS, LOINC, and RxNorm medication concepts. We compared Baseline, Hybrid biomedical retrieval-augmented generation (RAG), and terminology-guided Family expansion pipelines using Claude, GPT Qwen, and MedGemma on 40 criteria from 11 trial groups. Performance was evaluated against expert-built reference code sets using exact-code precision, recall, and F1. Results The optimal pipeline varied by model. Claude with Baseline achieved the highest performance (precision 0.755, recall 0.619, F1 0.680), followed by GPT-5.5 with Baseline (precision 0.569, recall 0.658, F1 0.610), Qwen with Hybrid biomedical RAG (precision 0.656, recall 0.470, F1 0.548), and MedGemma with Family expansion (precision 0.487, recall 0.316, F1 0.383). Hybrid biomedical RAG improved aggregate F1 only for Qwen but increased GPT-5.5 RxNorm F1 from 0.320 to 0.909. Macro-averaged results showed criterion-level gains despite lower micro-averaged aggregate performance. Family expansion increased recall across models but generally reduced precision. In staged verifier ablation, micro-F1 increased from 0.254 before verification to 0.505 after final verification and expansion. Existence/vocabulary checking removed 2,594 false-positive codes, and acceptance filtering removed 952 additional false-positive codes before controlled expansion. Conclusions Combining LLM-based clinical interpretation with deterministic terminology verification produces auditable, database-ready code sets, but retrieval and broad family expansion do not consistently improve exact-code performance. Retrieval was particularly useful for RxNorm mapping, whereas overly broad or incomplete candidate generation remained the main source of error. Deterministic verification improves code validity and query readiness but cannot replace accurate clinical interpretation.

15
Developing and Prospectively Validating a Reproducible Graph Representation Specification for Clinical Guideline Algorithms: The Measurement Foundation of the Clinical Guideline Complexity Index

Milani, R. V.; Bober, R. M.

2026-07-20 health informatics 10.64898/2026.07.17.26358358 medRxiv
Top 0.1%
34.0%
Show abstract

Background. Translating a clinical guideline decision algorithm into a computational graph requires judgment, and unconstrained coding yields divergent graphs; any complexity measure computed from such a graph inherits that variation, so its reproducibility must be demonstrated rather than assumed. Objective. To develop, and prospectively test, an empirical method for making graph extraction reproducible, using the Clinical Guideline Complexity Index (CGCI) and four guideline algorithms as a case study. Methods. We built a Graph Representation Specification (an ontology, a motif catalogue, disambiguation conventions, decomposition rules, a deterministic validator, and a scoring engine) and refined it by error-driven grammar induction: measure inter-coder disagreement, localize its dominant class, induce a single grammar rule, and prospectively test whether that rule improves agreement in the anticipated class. Reproducibility was quantified with a pre-specified, topology-based endpoint (Decision Topology Agreement) rather than edge agreement, which is oversensitive to representational choices that do not affect the score. Two trained coders independently coded the diabetes, dyslipidemia, heart-failure, and hypertension algorithms. Results. A rule induced from the diabetes comorbidity panel (assessment topology) generated a pre-specified prediction that heart-failure figures, sharing the same motif, would converge; on a fresh, independently coded pair they did, with an absolute CGCI difference of approximately one. Decision topology reproduced closely (decision-order agreement at or near 1.00 for three of four guidelines), while breadth counting was rule-sensitive: an explicit modifier-counting rule reduced the largest disagreement from 27 to 4 tokens. Residual disagreement was bounded and localizable to specific, nameable representational choices. Conclusions. Graph-extraction reproducibility can be systematically improved through iterative grammar refinement, and a prospectively derived rule can be confirmed to improve agreement. These results establish the measurement foundation (reliability, not construct validity) for a companion study interpreting CGCI as cognitive load, and the method may apply wherever graphs are extracted from structured source artifacts.

16
DBToken: A Database Tokenizer for Medical Event Foundation Models

Shin, I.; McCann, K.; Marino, G.; Siam, U. T.; Li, H.; Stutz, E.; Edara, R.; Loza, A. J.

2026-08-21 health informatics 10.64898/2026.08.18.26360487 medRxiv
Top 0.1%
33.2%
Show abstract

Objectives Transformer models for electronic health records require converting clinical data into token sequences, however standardized tokenization and evaluation frameworks are lacking. We introduce DBToken, an open-source library, and bits-per-row (BPR), a metric for comparing tokenization strategies. Materials and Methods DBToken accepts Medical Event Data Standard (MEDS)-compatible input and supports multiple text, numeric, and temporal tokenization strategies. BPR extends the bits-per-byte metric used in language models to enable comparison across tokenization strategies. Results DBToken efficiently tokenized data across configurations. BPR identified the vocabulary size associated with the best clinical outcome performance and localized differences in numeric tokenization performance by token class. Discussion Optimal tokenization strategies for medical foundation models are a subject of active research. DBToken enables reproducible tokenization experiments, while BPR efficiently screens vocabulary sizes and numeric representations before downstream evaluation. Conclusion DBToken and the BPR metric provide open-source infrastructure for reproducible EHR tokenization and cross-strategy evaluation.

17
Aggregating data to accelerate personalized therapy in heart failure (ADAPT-HF)

Roeder, C.; Goerg, C.; Talebi, A.; Stevens, L. M.; Scholtens, D. M.; Rasmussen-Torvik, L. P.; Alagna, L. M.; Shah, S. J.; Hall, J. L.; Das, A. K.; Jhund, P. S.; Kao, D. P.

2026-07-16 health informatics 10.64898/2026.07.13.26357501 medRxiv
Top 0.1%
33.2%
Show abstract

Background: Increased public access to data from disparate sources provides opportunities to study and validate predictive and subphenotype models in heterogeneous disease conditions using aggregated individual patient data. Robust, explicit, and transparent harmonization of data elements is critical to ensure interpretability, reproducibility, and generalizability of secondary and retrospective analyses. Methods & Results: We designed and implemented ADAPT (Aggregating Data to Accelerate Personalized Therapy), a scalable framework using multiple software packages (R, SQL, BigQuery) that enables rapid, explicit harmonization of structured data elements from randomized trials and observational studies using a standard spreadsheet interface. User-specified criteria are applied to primary study data to produce harmonized longitudinal datasets comprised of demographics, medical history, quantitative observations, repeated measures, and clinical outcomes. We demonstrate this functionality using 26 clinical studies found in the National Heart, Lung, and Blood Institute BioLINCC resource. We illustrate the scalability of ADAPT to the order of billions of datapoints using administrative clinical data in a cloud-computing platform. We also present examples of collaborators using ADAPT for independent harmonization tasks for secondary analyses and democratization of publicly available data. Conclusion: ADAPT is a disease-agnostic, extensible, and scalable platform to support robust, transparent harmonization of structured research data using interfaces accessible to a variety of researchers regardless of programming ability. It extends FAIR principles beyond research data to also represent harmonization analyses by improving Findability of harmonization decisions, Accessibility of methods to other stakeholders, Interoperability with independent analyses and datasets, and Reusability through efficient implementation in a variety of analysis environments.

18
FHIRBench: Benchmarking FHIR Clinical Data Serialization Strategies for Large Language Models

Chong, J.

2026-07-15 health informatics 10.64898/2026.07.14.26358020 medRxiv
Top 0.1%
32.7%
Show abstract

We present FHIRBench, a benchmark evaluating six FHIR clinical data serialization strategies across four frontier LLMs (Claude Sonnet 4.5, GPT-5.4, DeepSeek V3.2, Qwen3 32B) on three clinical tasks using 100 stratified synthetic FHIR R4 patient bundles. We employ two evaluation layers: token-level F1 and LLM-as-judge rubric on four clinical dimensions, yielding 7,200 evaluations per layer. Our findings reveal four results. First, serialization significantly impacts quality but the direction diverges between layers: Condensed outperforms Raw JSON on F1 for 3/4 models (Wilcoxon p < 10^-17), while Raw JSON achieves higher judge scores for 3/4 models (p < 10^-7). Narrative achieves 95% of Raw JSON's quality at 83% fewer tokens. Second, model rankings completely reverse between layers -- Claude ranks last on F1 but first on clinical quality (p = 1.0 x 10^-6), demonstrating that single-metric evaluation produces misleading model selection. Third, a significant Model x Serializer interaction (Friedman p = 0.0009) precludes universal format recommendations, with GPT-5.4 favoring Raw JSON while open-weight models favor compressed formats. Fourth, Llama 3.1 70B exhibits 100% inference failure on complex patients despite operating within its nominal context window, revealing a patient-safety gap where AI fails for the patients who need it most. These findings establish that clinical AI systems require model-aware serialization middleware, multi-layer evaluation frameworks, and capacity verification before deployment. Code and data publicly available.

19
Exploring the Application of the Observational Medical Outcomes Partnership Common Data Model to Multi-site Stroke Rehabilitation Research Data

Loomis, K. J.; Kumar, A.; Marin-Pardo, O.; Bellinger, G. C.; French, M. A.; Roemmich, R. T.; Liew, S.-L.

2026-07-08 health informatics 10.64898/2026.06.28.26356618 medRxiv
Top 0.1%
32.3%
Show abstract

Background: Emerging artificial intelligence and machine learning (AI/ML) tools can help generate robust knowledge to support precision rehabilitation approaches for varied patient populations. There is a large amount of research-generated and clinical rehabilitation data available for this purpose; however, a pronounced lack of interoperability prevents large-scale data aggregation. Common data models (CDMs) such as Observational Medical Outcomes Partnership (OMOP) have improved data interoperability across healthcare settings, and more recently, for clinical rehabilitation data, specifically. However, the application of these CDMs to research-generated data has not yet been explored. Therefore, as a foundational step, our study evaluated the breadth and depth of OMOP CDM coverage for data in a multi-site repository of harmonized rehabilitation research data: the Enhancing NeuroImaging Genetics through Meta-Analysis Stroke Recovery (ENIGMA-SR) database. Methods: Two raters independently mapped data elements representing 46 demographics and medical history (DMH) ENIGMA-SR variables and 95 distinct ENIGMA-SR rehabilitation assessments to OMOP standard concepts. Initial rater agreement was assessed for data element inclusion in OMOP and for specific OMOP concepts used (primary metric: Gwet's agreement coefficient [AC]). Mapping differences were reconciled, and final mappings were descriptively analyzed to examine (1) overall OMOP inclusion, (2) inclusion of more granular levels (subscales, items) of complex assessments, and (3) mapped OMOP concept characteristics. Results: Initial rater agreement was good/very good for overall OMOP inclusion of DMH and assessment data elements and for OMOP concepts mapped across almost all assessment data elements (Gwet's AC: 0.79-0.89). Initial OMOP concept agreement was more variable for DMH data elements; however, all mapping differences were successfully reconciled to 100%. Overall, DMH data elements had higher OMOP inclusion than rehabilitation assessments: 84.8% (39/46) vs. 58.9% (56/95). OMOP coverage was particularly limited for complex assessment subscale- and item-level data elements (9.4% [3/32]; 19.2% [14/73]) and did not match the granularity level represented in ENIGMA-SR data for 56.2% (41/73) of complex assessments. DMH and top-level assessment data elements were frequently mapped to multiple OMOP concepts (median: 6, 2; range: 1-23, 1-8), and for > 50% of these data elements the concepts spanned 2-3 different OMOP domains. Conclusion: For ENIGMA-SR, the OMOP CDM has good coverage of DMH data, moderate top-level coverage of rehabilitation assessments, and very limited coverage of assessment subscales and items. This uneven coverage, combined with variability in OMOP concepts and domains mapped to equivalent data points, presents challenges for aggregating clinical and research-generated rehabilitation data into AI/ML-ready datasets. Moreover, software tools currently available to facilitate the mapping process do not effectively accommodate content- and structure-related features inherent to research-generated data. Going forward, the utility of the OMOP CDM to aggregate multi-source rehabilitation data may be improved by expanding the catalogue of OMOP rehabilitation-related concepts, building cross-walks to research-oriented data standards, and adapting emerging computational tools to streamline the mapping process.

20
VarEx: A Large Language Model Pipeline for Automated Extraction of Exposures, Outcomes, and Covariates from Epidemiologic Studies

Malec, S. A.; Pradhan, M.; Upadhayaya, R.; Metzger, V.

2026-06-15 health informatics 10.64898/2026.06.13.26355589 medRxiv
Top 0.1%
31.2%
Show abstract

Objective: Observational studies are essential for investigating risk factors for Alzheimer's disease and related dementias (ADRD), but inconsistent reporting and selection of covariates can contribute to residual confounding, omitted-variable bias, and reduced reproducibility. We developed and evaluated VAREX (Variable Extraction), a large language model (LLM)-based information extraction framework designed to automatically identify exposures, outcomes, and covariates from epidemiologic studies and populate structured evidence repositories. Materials and Methods: VAREX combines retrieval-augmented generation, biomedical language-model embeddings, semantic chunking, cross-encoder reranking, and prompt-engineered LLM workflows to extract epidemiologic variables from full-text biomedical articles. The framework was evaluated using a reference-standard corpus of observational studies examining blood pressure variability (BPV) and Alzheimer's disease-related dementias (ADRD), together with external validation datasets involving other exposure-outcome relationships. Extracted variables were compared with independently curated human reference standards using semantic matching and one-to-one assignment procedures. Covariates were additionally classified into ten epidemiologically relevant semantic categories. Results: In the primary BPV[-&gt;]ADRD corpus (10 studies), VAREX achieved a precision of 0.91, recall of 0.84, and F1-score of 0.87 for variable extraction. Covariate classification accuracy was 0.90, yielding a strict extraction-and-classification F1-score of 0.78. External validation datasets demonstrated comparable performance across diverse epidemiologic domains, with extraction F1-scores ranging from 0.73 to 0.85. Category-level performance was strongest for health behaviors (F1=0.96), sociodemographic variables (F1=0.90), and medication exposures (F1=0.89). Compared with published estimates of manual systematic-review effort, VAREX reduced processing time from approximately 61 minutes to 9 minutes per article, representing an 85.7% reduction in review time. Discussion: These findings demonstrate that LLM-based information extraction can accurately identify and classify epidemiologic variables across heterogeneous observational-study designs. Automated extraction enables scalable construction of structured repositories of exposures, outcomes, and covariates while substantially reducing the labor required for evidence synthesis and systematic reviews. Conclusion: VAREX provides an effective framework for automated extraction and classification of epidemiologic variables from the biomedical literature. By supporting large-scale evidence synthesis and structured knowledge resource development, VAREX may facilitate more rigorous observational research, improved confounder identification, and enhanced reproducibility in epidemiology.